Papers with South Asian
Bhaasha, Bhāṣā, Zaban: A Survey for Low-Resourced Languages in South Asia – Current Stage and Challenges (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a survey examines the current efforts and challenges of NLP models for South Asian languages . there are more than 650 languages in South Asia, but many have very limited computational resources or are missing from existing models. |
| Approach: | a survey examines efforts and challenges of NLP for South Asian languages . they focus on transformer-based models such as BERT, T5, & GPT . findings highlight substantial issues, including missing data in critical domains . |
| Outcome: | The findings highlight significant issues, including missing data in critical domains . the survey aims to raise awareness within the NLP community for more targeted data curation . |
Processing South Asian Languages Written in the Latin Script: the Dakshina Dataset (2020.lrec-1)
Copied to clipboard
Brian Roark, Lawrence Wolf-Sonkin, Christo Kirov, Sabrina J. Mielke, Cibu Johny, Isin Demirsahin, Keith Hall
| Challenge: | a new resource is available for 12 South Asian languages that use the Latin script for text entry . the Latin-script system is not widely used in South Asian language writing, despite the Latin alphabet . |
| Approach: | They describe the Dakshina dataset, a new resource consisting of text in both the Latin and native scripts for 12 South Asian languages. |
| Outcome: | The Dakshina dataset includes text in both the Latin and native scripts for 12 languages . the authors provide baseline results on several tasks made possible by the dataset . |
Multilingual Coreference Resolution in Low-resource South Asian Languages (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing coreference resolution models for South Asian languages are limited . a a sanity check for the prediction of translations is required to ensure accuracy of the model, authors say . |
| Approach: | They evaluate an end-to-end coreference resolution model on a Hindi golden set . they use translation and word-alignment tools to translate a translated dataset into 31 languages . |
| Outcome: | The proposed model scored 64 and 68 on a Hindi golden set. |